Skip to content

rl: add the elastic resource benchmark contract, gating, evidence and paired work - #66

Merged
michaellchung merged 1 commit into
mainfrom
feat/rl-elastic-benchmark
Sep 29, 2026
Merged

michaellchung merged 1 commit into
mainfrom
feat/rl-elastic-benchmark

Conversation

@michaellchung

Copy link
Copy Markdown
Contributor

The GPU-free half of the rl-elastic-resource-benchmark study harness: scripts/benchmark_rl_elastic.py plans and reports studies comparing fixed GPU partitions (trainer / rollout / standby) against scheduled and automatic resizing inside one RL island.

What lands

  • Versioned study manifest with a canonical hash, in calibration and formal modes
  • Config / edge / pool validation gated behind runtime capability attestation — an item only becomes runnable once a runner attests it can run it
  • Evidence index with tamper-rejecting resume
  • Dry-run plan CLI, layered results and a report that exits non-zero while the matrix is incomplete
  • Paired splits, schedules and scenarios reusing the legacy benchmark_rl pairing through yeto/rl/elastic_benchmark/legacy.py

scripts/benchmark_rl.py, its CLI and its result fields are unchanged.

No GPU runner ships here. None of the four commands loads a model, imports torch/ray/miles, or creates cloud resources.

Conflicts

None — every file is new (yeto/rl/elastic_benchmark/, the script, the test module, the doc). Nothing existing is touched.

Spec

The planning contract lives in the miles repository under openspec/changes/rl-elastic-resource-benchmark/, as docs/RL_ELASTIC_BENCHMARK.md states; there is deliberately no OpenSpec change for it in this repo.

Tests

19 pass locally (tests/test_rl_elastic_benchmark.py).

🤖 Generated with Claude Code

…red work

Implements the GPU-free half of the rl-elastic-resource-benchmark change:
versioned study manifest with canonical hash and calibration/formal modes,
config/edge/pool validation gated by runtime capability attestation,
evidence index with tamper-rejecting resume, dry-run plan CLI, layered
results and report, and paired splits/schedules/scenarios that reuse the
legacy benchmark_rl pairing without changing its CLI.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@michaellchung
michaellchung merged commit 8653152 into main Sep 29, 2026
1 of 6 checks passed
michaellchung added a commit to michaellchung/yeto that referenced this pull request Sep 29, 2026
…ed fingerprint, pause audit fixes)

- tasks 1.4/1.5/1.6 unchecked: implemented + CPU tests pass, dependencies
  (1.2/1.3, GPU X9) not met; progress.md updated.
- accept spec spelling "serial-colocated" as alias of "colocated-serial"
  across manifest, attestation, EngineCapabilities and ExecutionProfile.
- fingerprint_rejection fails closed when the study fingerprint is None or
  "unresolved"; agentenv#66 tests pin a fingerprint where they expect support.
- pause-audit.md: read_loop line 1259, decoupled budget mode
  max_reconnects=None, BUDGET_DONE lease only fatal in the
  collect_budget_reports window in learner-budget mode, non-budget decoupled
  marked unproven, 450 s documented as policy not a code limit.
- pause_audit.py: budget derived from quorum_timeout_s x margin; unused
  constants removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
michaellchung added a commit to michaellchung/yeto that referenced this pull request Sep 29, 2026
…alized (Verda not supported), DynaResize hypotheses with page citations

runtime_manifest.py collects commits/import paths/versions/fork interfaces in the
image and refuses certification on pin mismatch or a missing interface for a
declared capability. gpu-plan.md section 8 maps experiments to the existing
launcher (Nebius sky, Modal runner), agentenv#66 pool identity, per-rental record and
mechanical cleanup; notes the Modal H100 (no '!') launcher gap. DynaResize
(arXiv:2607.22614) turned into H1-H6 with page numbers, miles adaptation
points and negative-result handling; no paper constants adopted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant